Frontiers in Artificial Intelligence
○ Frontiers Media SA
Preprints posted in the last 30 days, ranked by how well they match Frontiers in Artificial Intelligence's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Darko, R.; Dwumah, D.; Agyapong, K. S.; Agyenim-Boateng, Y.; Darko Anim, R.; Wisdom Jakper, J.; Owusu-Ansah, N. K.; Owusu-Ansah, R.
Show abstract
Machine learning workflows frequently incorporate data preprocessing to enhance predictive performance. However, the need for Super Learner ensembles made up only of preprocessing-invariant tree-based algorithms remains unexplored. Using three benchmark clinical classification datasets, this study examined how preprocessing affected the Super Learner's prediction performance, learner weight distribution, and oracle behavior. The Heart Disease (207 observations), Indian Liver Patient Dataset (583 observations), and Pima Indians Diabetes (768 observations) datasets were used to create a Super Learner ensemble model that included Classification and Regression Trees (CART), Random Forest, Ranger, and Extreme Gradient Boosting (XGBoost). Models were evaluated under raw and preprocessed data conditions using repeated cross-validation. Predictive performance was assessed using the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), and Brier score. Learner weight allocation and Oracle Gap were compared using paired Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Preprocessing produced negligible changes in predictive performance for the Heart Disease and Pima datasets. For the ILPD dataset, preprocessing significantly improved AUC (0.746 to 0.752; adjusted p = 0.0017) and reduced the Brier score (0.177 to 0.175; adjusted p < 0.001). Learner weights remained largely stable, although Random Forest replaced Ranger as the dominant learner for the Heart Disease dataset. Oracle Gaps remained extremely small (<0.002) across all datasets and did not differ significantly between preprocessing conditions. Preprocessing provides limited benefit for Super Learner ensembles composed of preprocessing-invariant learners and does not materially alter their oracle behavior. Preprocessing decisions should therefore be guided by dataset characteristics rather than adopted as a universal modelling practice.
Pybus, A.; Qiu, J.; Morais Lyra, P. C.; Dang, K.; Narvaez-Bandera, I.; Jolaogun, T.; Goecks, J.
Show abstract
Survival analysis is a fundamental technique in biomedical research for modeling time-to-event data. It enables the identification of prognostic factors in disease, compares survival outcomes across treatment groups, and performs targeted treatment selection. A variety of machine learning (ML) approaches to survival analysis have emerged to complement classical statistical methods, especially for high-dimensional datasets with complex, nonlinear interactions between features. However, using survival ML methods requires addressing challenges such as censoring-unaware evaluation, overfitting, selecting performance metrics, and data leakage. To address these and other difficulties in using survival ML models, we developed the mlsurv software package. mlsurv is an open-source Python package built around three major design principles: 1) methodological rigor, including evidence-based model selection, leakage-free pipelines, and multi-metric evaluation, 2) multi-scale evaluation and interpretation, including population and subpopulation evaluation, patient-level explanations, and feature analysis, and 3) automated trust and transparency, including limitation flagging and TRIPOD+AI-aligned reporting. mlsurv bundles ten models spanning linear, ensemble, kernel, and deep learning families within a unified software package. We demonstrate mlsurv on the Chowell immunotherapy cohort (n=1,479). The survival-trained models achieve a test concordance index of 0.73 for overall survival prediction. Further, risk scores strongly correlate with the response-trained LORIS clinical score (|{rho}| up to 0.84), reflecting the overlap between prognostic and predictive signal. mlsurv enables biomedical researchers to conduct rigorous, multi-model survival analysis and benchmarking using minimal code with default best practices rather than implementing custom scripts and methodological safeguards from scratch.
Feng, W.; Liu, S.; Yang, Z.; Tao, Y.; Gu, X.; Jin, W.
Show abstract
Background Hepatocellular carcinoma (HCC) treatment selection demands nuanced integration of heterogeneous patient data, yet prevailing predictive models rely on restricted data modalities and oversimplified therapeutic frameworks, compromising clinical translation. Objective We developed and validated a multimodal artificial intelligence framework to guide optimal treatment strategy selection across the full spectrum of HCC interventions. Methods This retrospective study comprised 1,043 HCC patients (development cohort, January 2017-December 2023) and 55 external validation patients (2023) from Wuxi Peoples Hospital. We engineered Embedding-Augmented Extra Trees (ET-Emb), a novel model fusing structured clinical variables with contextual text embeddings derived from medical histories and radiology reports. ET-Emb quantifies probabilities for five primary treatments: open/laparoscopic resection, transarterial chemoembolization, radiofrequency ablation (RFA), and chemotherapy. Model performance was rigorously assessed via 10-fold cross-validation and external validation using ROC-AUC and PR-AUC metrics. Results ET-Emb demonstrated robust performance in the development cohort (ROC-AUC: 0.84 {+/-} 0.04; PR-AUC: 0.55 {+/-} 0.06), significantly outperforming established benchmarks. This generalizability was preserved in external validation (ROC-AUC: 0.77 {+/-} 0.02; PR-AUC: 0.47 {+/-} 0.03). SHAP analysis identified textual clinical narratives and socioeconomic determinants as critical predictive drivers. Conclusions By unifying structured and unstructured data modalities, ET-Emb delivers accurate, multi-treatment strategy prediction for HCC. Its clinical validity and the demonstrated significance of textual features establish multimodal AI as an essential paradigm for simulating complex oncological decision-making, positioning ET-Emb as a transformative tool for precision HCC management.
Ritter, M.; Bogadhi, A. R.
Show abstract
"Revealing the structure of pharmacobehavioral space through motion sequencing" by Wiltschko et al. (2020) has been highly influential in behavioral phenotyping research. In a cohort of nearly 700 mice, the authors demonstrated that Motion Sequencing (MoSeq) could distinguish behavioral effects across a large and diverse set of neuroactive and psychoactive compounds. A central conclusion of the study is that MoSeq syllable features substantially outperform more traditional scalar behavioral features in treatment classification tasks. Although this comparison is not emphasized outside the Results section, the reported advantage corresponds to an increase in classification performance exceeding 50% relative to scalar feature representations. While reproducing parts of the analysis using the publicly available dataset, we found that much of this apparent performance difference can be attributed to differences in preprocessing, classifier selection, and hyperparameter optimization. Under alternative, but comparably standard, analytical choices, the performance gap between scalar features and MoSeq syllables was reduced to approximately 11%. Furthermore, in our reanalysis, the performance advantage of MoSeq syllables became statistically significant primarily in highly dense pharmacobehavioral spaces. These findings do not contradict the utility of MoSeq syllables. Rather, they suggest that the magnitude and generality of their advantage over simpler scalar features may depend strongly on analytical methodology and dataset structure. This distinction is practically relevant, as scalar feature approaches are substantially less computationally demanding and often easier to interpret biologically. Consequently, for laboratories with limited computational resources or for studies focused on specific treatment effects, conventional scalar representations may provide a competitive and more accessible alternative. Our findings highlight the importance of analytical standardization and reproducibility in comparative behavioral representation studies.
Bisaso, K. R.; Kadada, K. R.; Bisaso, K. S.; Ette, E. I.
Show abstract
Background: Parametric time-to-event models require specification of a baseline hazard function, which may influence prediction when the underlying hazard shape is uncertain. This study compared conventional joint longitudinal time-to-event models with mechanistic Multi-Task Logistic Regression, which directly models the survival distribution without selecting a continuous parametric hazard family. Methods: A simulated dataset of 100 individuals with longitudinal sum of longest diameters and event outcomes was analyzed using a shared mechanistic tumor shrinkage regrowth model. Event submodels comprised exponential, Gompertz, Weibull, log-normal, log-logistic, and circadian hazards, mechanistic Multi-Task Logistic Regression, and a hybrid neural-mechanistic extension. All models were estimated jointly using shared patient-specific random effects and longitudinal data. Models were evaluated using longitudinal goodness-of-fit, visual predictive checks, five-fold cross-validated inverse-probability-of-censoring-weighted dynamic area under the curve and Brier scores, integrated Brier score, calibration, and event-interval negative log score. Results: Longitudinal parameter estimates and diagnostics were comparable across models. All conventional hazard models produced identical dynamic area under the curve values within prediction windows, although probabilistic accuracy differed. The log-normal hazard achieved the lowest overall integrated Brier score (0.1928). Mechanistic Multi-Task Logistic Regression achieved the highest later landmark discrimination (area under the curve 0.867 versus 0.798 for all hazard models) and the lowest mean event-interval negative log score (2.362). The hybrid model improved intermediate-landmark discrimination but not overall probabilistic accuracy. Conclusions: Mechanistic Multi-Task Logistic Regression provided competitive joint time-to-event prediction while avoiding baseline hazard-family selection. It represents a practical complementary alternative to parametric hazard modeling, particularly when hazard shape is uncertain and dynamic discrimination is important.
Ngamsaowaros, T.; Bodala, I.; Michopoulou, S.; Niranjan, M.
Show abstract
Predicting the course of Alzheimer's disease for individual patients remains a major challenge due to the heterogeneity of disease expression and the sparsity of longitudinal data. We introduce a variational Disease Progression Score (DPS) framework that maps multimodal biomarker dynamics (Cerebrospinal fluid, neuroimaging, and cognitive assessments) onto a continuous latent timeline with quantified uncertainty. The framework combines a neural encoder, which infers subject-specific progression parameters from demographic and clinical features, with a cascade of logistic functions structured according to the amyloid cascade hypothesis. Applied to the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort, the inferred timeline separated diagnostic groups it never observed (AUC 0.98 for cognitively normal vs Alzheimer's Disease), and the estimated cascade strengths and biomarker orderings were consistent with the established sequence of Alzheimer's pathology. The model produces individualised prognoses for previously unseen subjects from baseline data alone, with 95\% credible intervals achieving 89-98\% empirical coverage across biomarkers, and these predictions can be dynamically refined as new observations become available. The framework thus provides a biologically interpretable, uncertainty-aware index of disease severity, offering a probabilistic foundation for patient-level prognosis and precision monitoring in Alzheimer's disease.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
Nastaro, C. D.; Correa, B. R.; Tarantini, G.; Marana, S. R.; Cafe Ferreira, R. d. C.
Show abstract
Active teaching methodologies have been widely used to promote meaningful learning and student autonomy. In this context, quantitative approaches can help assess how students organize and integrate knowledge throughout the learning process. Among these approaches, semantic co-occurrence networks stand out, as they are capable of identifying relationships between words and revealing the conceptual structure of textual productions. The objective of this study was to investigate whether semantic network analyses can characterize differences in students conceptual organization in Microbiology during their participation in the active teaching methodology "Adopt a Bacterium." To this end, a case study was conducted in the Bacteriology course at the Institute of Biomedical Sciences of the University of Sao Paulo, analyzing the textual productions of two groups of students in the years 2024 and 2025 during their study of the bacterial genus Bacillus. The texts were evaluated using semantic co-occurrence networks, taking into account metrics of structure and conceptual integration. The results showed that both groups covered the microbiological content outlined in the course, though with different thematic focuses and approaches to integrating the concepts. Although both years featured modular structures (a statistical mode of 9 subgraphs), in 2025 the network exhibited greater discursive robustness (2 to 4 times more words with high Betweenness centrality) than in 2024. It is concluded that semantic network analysis allows for the characterization of differences in conceptual organization among students using active learning methodologies, serving as a complementary tool for assessing meaningful learning in Microbiology.
bolin, k.; Stibrant Sunnerhagen, K.
Show abstract
Background The time trend in long-term survival after a stroke is to some extent unknow due to (relatively) short follow up periods in available data. The objective of this study is to identify and quantify differences in long-term stroke survival in Sweden between men and women and patients with different attained educational levels, comparing two time-periods, 2000-2009 and 2010-2022. Methods This study employs total population Swedish register data pertaining to hospital-based care and mortality due to stroke for the period 2000-2022 in order to estimate survival (all-cause mortality) after ischaemic and haemorrhagic stroke, respectively, and pertaining to attained educational level. Kaplan-Meier survival functions are estimated stratifying for time-period, sex and educational level. Cox regressions are employed to quantify mortality hazard ratios between the strata. Age is taken into account in complementary analyses (supplement). Results Taking only time-period (2000-2009 vs 2010-2022) into account resulted in significantly higher survival in the second period for ischaemic stroke patients (HR: 0.84; 95% CI: 0.83-0.84), while no significant difference could be detected for haemorrhagic stroke. Stratifying for sex showed that men gained more than women in terms of reduced mortality hazard rate between the periods. Further stratifying by educational level and estimating survival separately for men and women showed that, for both men and women, patients with the lowest education were relatively worse off (compared to patients with higher education) in the second period. Further analyses, taking age into account, reversed the relative hazard ratio between men and women, but corroborated the result that low education is associated with poorer outcome than high education. Conclusions The results suggest that there are considerable differences in expected long-term survival after stroke between the sexes, but that this may be due to differences in age between the sexes at the time of stroke. Moreover, lower educational level is significantly associated with lower long-time survival.
Maleki, C.; Bertrand, Y.; Gailly, F.
Show abstract
Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sensitivity analysis. The framework is evaluated using an NHANES-derived fasting cohort for classification of documented diabetes status. The full fasting analysis cohort contained 2,582 participants, and a non-diagnostic laboratory subgroup, Gate0, contained 2,111 participants. On untouched test data, the rule-ensemble model achieved ROC-AUC and PR-AUC values of 0.959 and 0.873 in the full fasting cohort and 0.861 and 0.499 in Gate0. Four clinically interpretable candidate rules were selected using validation data only. A nonnegative survey-weighted logistic model removed one redundant rule and converted the remaining three binary activations into an auditable DMN score and model-estimated probability. The final DMN achieved ROC-AUC 0.769, PR-AUC 0.153, and Brier score 0.029 in the untouched Gate0 test set. In small rule-defined test subgroups, hypothetical five-unit BMI reductions lowered mean model-estimated probability by 2.40 to 5.89 percentage points when one or more BMI thresholds were crossed. These findings characterize policy sensitivity rather than causal effects and require external validation.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Dashti, N.; Schneider, M. M. K.; Eckardt, J. N.; Fiebig, F.; Schweigler, D.; Buttner, S.; Middeke, J. M.; Bornhauser, M.; Rollig, C.; Kather, J. N.; Wiest, I. C.
Show abstract
Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standardized MedDRA (Medical Dictionary for Regulatory Activities) coding. However, manual Low-Level Term (LLT) assignment remains labor-intensive, subjective, and difficult to scale. Although large language models (LLMs) have emerged as promising decision-support tools for automated coding, unguided zero-shot generation remains insufficient for reliable fine-grained MedDRA coding. Objective: To develop and evaluate a retrieval-augmented reasoning pipeline for clinically aligned LLT-level MedDRA coding of free-text adverse events from prospective AML clinical trials. Methods: We implemented a retrieval-augmented reasoning pipeline inspired by the retrieval-augmented generation (RAG) paradigm using LLaMA-3.3-70B-Instruct as the primary backbone and benchmarked the framework across multiple open instruction-tuned LLMs. Dense semantic retrieval first generated a constrained top-100 LLT candidate set for each AE, followed by structured LLM reasoning to select a single best-matching LLT and deterministic mapping to Preferred Term (PT) and System Organ Class (SOC) levels. The pipeline was evaluated retrospectively on AE datasets from three prospective AML clinical trials (MOSAIC, DELTA, and DaunoDouble) with automated LLT/PT/SOC metrics and expert-assessed Clinical Correctness Rate (CCR). Results: Clinical expert review showed high clinical acceptability of the RAG pipeline across datasets (91-97%). Under automated evaluation, the pipeline achieved LLT exact accuracy of 50-58%, PT accuracy of 78-85%, and SOC accuracy of 90-93%. Zero-shot generation and random candidate selection performed substantially worse. Semantic retrieval more often included the coder-assigned LLT among the candidate terms available to the model than retrieval based on lexical similarity. Multi-model benchmarking showed that backbone choice mainly affected LLT exact agreement, whereas PT and SOC performance remained comparatively stable. Conclusions: Retrieval-augmented reasoning supports clinically aligned MedDRA coding of free-text adverse events under realistic candidate constraints in AML clinical trials. Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.
Sadia, H.; Doyon, N.; Duchesne, S.
Show abstract
Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Bagchi, R.; Yee, N. J.; Kwon, J. Y.; Taseh, A.; Ashkani-Esfahani, S.
Show abstract
Purpose To evaluate whether domain-adaptive self-supervised pretraining on musculoskeletal radiographs improves fracture classification and attribution faithfulness relative to ImageNet-pretrained baselines. Materials and Methods This study (June 2025 to May 2026) used previously acquired radiographs to compare three ResNet-50 initializations: supervised ImageNet pretraining (control), self-supervised ImageNet pretraining (DINO), and DINO with additional domain-adapted pretraining on 44,029 musculoskeletal radiographs (DINO-Ortho). All models underwent supervised fine-tuning in three experiments: in-distribution (MURA and FracAtlas datasets), out-of-distribution (an external dataset of 5,365 calcaneal radiographs from 1,775 patients), and initial weights (calcaneal radiographs only). Metrics included sensitivity, specificity, test accuracy, area under the receiver operating characteristic curve (AUROC), and Cohen's kappa; attribution faithfulness was quantified using Remove and Debias scores from Grad-CAM saliency maps. Comparisons used DeLong and Friedman tests. Results Classification performance did not differ significantly between DINO-Ortho and either baseline in any experiment (DINO-Ortho AUROC, 0.89 in-distribution and 0.95 with initial weights). All three models discriminated poorly out-of-distribution (control, 0.59; DINO, 0.57; DINO-Ortho, 0.58). DINO-Ortho showed significantly higher attribution faithfulness than both baselines in all three experiments, including out-of-distribution (25.39 vs -10.41 and 2.14; P < .001) and initial weights (20.88 vs 11.51 and 1.27; P < .001). Qualitative rankings favored DINO-Ortho but did not differ significantly. Conclusion Domain-adapted self-supervised pretraining on musculoskeletal radiographs improved attribution faithfulness while maintaining classification performance comparable to ImageNet-pretrained baselines; no model generalized adequately to external radiographs without task-specific fine-tuning.
Turon, R.; Reining, L. C.; Hummel, P. A.; Schmittwilken, L.; Lind, C.; Yu, A. J.; Rothkopf, C. A.; Jaekel, F.; Wallis, T. S. A.
Show abstract
Behavioral experiments are often infeasible when stimulus spaces have many dimensions or when testing time is limited. One way to address this challenge is adaptive stimulus selection, where informative stimuli are chosen dynamically based on participants responses. However, in high-dimensional spaces, identifying such stimuli is computationally demanding. Here, we describe High-dimensional Online Particle Estimation (HOPE), which selects informative stimuli in less than a second for up to 50 dimensions, enabling efficient estimation of high-dimensional psychometric functions. We validate HOPE through simulations and a face-categorization experiment in an 18-dimensional parameter space with human participants. Compared to uniform stimulus presentation, HOPE reduces uncertainty over model parameters two-to three-times faster, reaching the same certainty in half the trials or fewer. This efficiency enables psychophysical studies that were previously impractical due to the exponential scaling of trial requirements.
Margolis, S. J.; Maddipatla, N. V. S. K.; Ioannides, K. L. H.; Wisk, L. E.; Schriger, D. L.; Elmore, J. G.
Show abstract
We evaluated nine patient-facing artificial intelligence products using 60 physician-developed standardized clinical cases and 540 multi-turn simulated patient encounters. Although overall triage accuracy showed no statistically significant difference across product categories, referral behavior differed substantially. Branded health AI products more frequently over-triaged low-acuity cases (28% vs 3% vs 2%) and recommended affiliated, fee-requiring clinical services. These findings suggest evaluation of patient-facing medical AI should assess referral behavior alongside overall triage accuracy.
Adapa, K.; Mosaly, P. R.; Yu, F.; Moore, C.; McGurk, R.; Das, S.; Mazur, L.
Show abstract
Radiation oncology has a long history of developing in-house health information technology (HIT) tools such as quality assurance (QA) checklists, yet there is little guidance from professional bodies on how to implement these tools in complex clinical environments. Building on our previous work that used human-centered participatory co-design, the Task-User-Representation-Function (TURF) framework, and multi-method usability evaluations to design and develop an enhanced dosimetry QA checklist (DQC), this study investigated the barriers and facilitators (determinants) to implementing the enhanced DQC in a radiation oncology clinic, examined implementation strategies, proposed an implementation framework for QA checklists in radiation oncology, and assessed four implementation outcomes: acceptability, appropriateness, feasibility, and adoption. We conducted a qualitative implementation study using an abductive research approach at an academic medical center. All key stakeholders (dosimetrists, physicists, trainees, and software developers) participated in semi-structured interviews, field observations, and surveys across pre-implementation, implementation, and post-implementation phases. Data were analyzed using a hybrid inductive-deductive approach, with deductive coding guided by an adapted Consolidated Framework for Implementation Research (CFIR) mapped to the Unified Theory of Acceptance and Use of Technology and by the Expert Recommendations for Implementing Change (ERIC) compilation. We identified 4 CFIR constructs and 12 sub-constructs as barriers, with structural characteristics and planning showing the highest negative valence, and 5 CFIR constructs and 19 sub-constructs as facilitators, with relative advantage, culture, and leadership engagement showing the highest positive valence. Participants' suggestions mapped to 19 ERIC strategies in 7 clusters, and the CFIR-ERIC matching tool identified 14 evidence-based strategies in 4 clusters that informed a proposed phased implementation framework. Acceptability, appropriateness, and feasibility scores improved significantly from pre-implementation to implementation for all professional roles (p<0.05), yet adoption reached 100% only in the sixth week of implementation. These findings highlight the value of combining subjective and objective implementation outcomes and provide a practical, evidence-based framework for implementing in-house QA checklists in radiation oncology that warrants validation in diverse settings.
Hui, J.; Xia, M.; Wilson, J.; Hill, E. D.; Scheer, A.; Franz, L.; Engelhard, M. M.; Goldstein, B. A.
Show abstract
The performance of an EHR-based deep learning model trained on a small sample can be improved if more data is collected. Instead of collecting more data, the model can be trained on additional data from an analogous external source. However, this risks the model learning patterns in the external data that do not generalize to the target sample. Furthermore, data use agreements often prohibit combining datasets with medical records of different sources. We consider utilizing pre-existing methods in continual learning, namely the elastic weight consolidation (EWC) loss function and variational continual learning (VCL), both of which are regularization-based methods that we use to borrow external data and incorporate parameters from a model on external data into local model training. To investigate the utility of this modeling framework, we consider two binary classification tasks: (1) predicting which children will be diagnosed with autism spectrum disorder (ASD) from medical claims up to 18 months, and (2) predicting which patients with end-stage renal disease (ESRD) will be re-hospitalized within 30 days. Target datasets were derived from Duke University's EHR warehouse, and external datasets were sourced from either NC Medicaid claims for the ASD prediction task, or the United States Renal Data System (USRDS) for the rehospitalization prediction task. For both of these tasks, borrowing models - using either the EWC loss function or VCL - performed similarly to that of a model trained only on the full external data, when the sample size of target data used to train the model was small. That is, while a model that does not borrow using our methods performed poorly in low data regimes, the borrowing model instead matched the performance of a model trained on external data even when sample size of target data was small. In addition, an analysis of model predictions showed that models with small samples are better calibrated and more functionally similar to a model trained only on external data when the sample size is small.
Champeaux, S. A.; Booth, J.; Brown, A.; Sebire, N. J.; Drobnjak, I.; Bowyer, S.
Show abstract
Background: Machine learning models leveraging electronic health records (EHRs) can support earlier detection of sepsis in intensive care units (ICUs). However, their clinical utility depends on reproducibility across institutions and patient populations. Building on a published pipeline from the Children's Hospital of Philadelphia (CHOP), this study examines how a neonatal sepsis prediction framework performs and can be adapted to a range of intensive care environments, paediatric, cardiac, and neonatal, at Great Ormond Street Hospital (GOSH). Methods: We extracted de-identified ICU EHR data from GOSH and applied feature derivation, unit harmonisation, and temporal sampling to align with the CHOP dataset used by Masino et al. (2019). Seven classifiers were first evaluated using CHOP-trained weights to characterise cross-domain behaviour and then retrained on local data to assess recoverability and site-specific adaptation. Model discrimination was summarised by AUC and F1, and learning curves were used to explore sample efficiency and bias-variance dynamics. Results: Models achieved strong discrimination on the CHOP neonatal cohort but demonstrated reduced performance when transferred to the mixed GOSH ICU population, reflecting anticipated domain and population shift. Retraining on GOSH data restored discrimination (AUC range 0.69-0.86), with Gradient Boosting (AUC 0.86 vs AUC 0.87 at CHOP) and KNN (AUC 0.80 vs AUC 0.79 at CHOP) models performing comparably to their CHOP benchmarks. DeLong's test confirmed statistically significant gains across all classifiers (p < 0.001). Conclusion: ICU cohort and baseline demographic differences between CHOP and GOSH introduced domain shift that limited direct model transfer. Elements of the original preprocessing pipeline could not be reproduced, further constraining transportability. Yet, retraining on local data restored high discrimination, showing that the modelling framework remains robust when re-estimated in new settings. These results highlight local adaptation as a practical route to recover performance and support safe, generalisable deployment of clinical prediction models in mixed clinical environments.